Papers by Janet B. Pierrehumbert
Stories that (are) Move(d by) Markets: A Causal Exploration of Market Shocks and Semantic Shifts across Different Partisan Groups (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing attempts to model the relationship between the real world and written or spoken text have focused on more interpretable and simplistic text representations. |
| Approach: | They propose to link shifts in semantic embedding space to real-world market shocks and partisanship to shape predictions of market fluctuations. |
| Outcome: | The proposed model demonstrates that partisanship can influence the predictive power of text for market fluctuations and shape reactions to those same shocks. |
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks (2025.acl-long)
Copied to clipboard
Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann, Adrian de Wynter, Xun Wang, Si-Qing Chen, Michael J. Wooldridge, Janet B. Pierrehumbert, Furu Wei
| Challenge: | a study aims to assess the fairness and robustness of Large Language Models in dialectal queries . speakers of "non-standard" dialects are known to experience implicit and explicit discrimination . |
| Approach: | They propose to use a benchmark to assess the fairness of large language models in dialects . they hire speakers with computer science backgrounds to rewrite seven popular benchmarks based on AAVE . |
| Outcome: | The proposed benchmarks show that most models show significant brittleness and unfairness to queries in AAVE. |
ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts (2025.emnlp-main)
Copied to clipboard
| Challenge: | Scientific fact-checking has largely focused on textual and tabular sources, neglecting scientific charts. |
| Approach: | They propose a benchmark for scientific fact-checking grounded in scientific charts . climateViz comprises 49,862 claims paired with 2,896 visualizations . results show current models struggle to perform fact- checking when statistical reasoning is required . |
| Outcome: | The climateviz benchmark is the first large-scale benchmark for scientific fact-checking . it includes 49,862 claims paired with 2,896 visualizations labeled as support, refute, or not enough . |
Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics (2024.lrec-main)
Copied to clipboard
| Challenge: | Scalar adjectives describe different domain scales and vary in intensity . they can be triggered by scalar adjective and require listeners to reason pragmatically about them. |
| Approach: | They probe different families of Large Language Models for their knowledge of the lexical semantics of scalar adjectives and one specific aspect of their pragmatics. |
| Outcome: | The proposed models encode rich lexical-semantic information about scalar adjectives but lack a good understanding of skalar diversity. |
Quantifying Compositionality of Classic and State-of-the-Art Embeddings (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Static word embeddings make strong claims about compositionality, but the SOTA generative models go too far in the other direction. |
| Approach: | a new study evaluates the compositionality of word embeddings by canonical correlation analysis . strong compositional signals are observed in later training stages across data modalities . |
| Outcome: | a new evaluation of compositional models shows that they exploit access meanings when justified . strong compositional signals are observed in later training stages and in deeper layers of the transformer-based model before a decline at the top layer. |
Actors, Frames and Arguments: A Multi-Decade Computational Analysis of Climate Discourse in Financial News using Large Language Models (2026.findings-eacl)
Copied to clipboard
| Challenge: | a new study examines how financial news media portrays climate change . financial news is the nervous system of the global economy . |
| Approach: | They propose a three-stage Actor–Frame–Argument pipeline that uses large language models to extract actors, stances, frames, and argumentative structures from a 980,061-article corpus. |
| Outcome: | The proposed pipeline extracts actors, stances, frames, and argumentative structures from a 980,061-article corpus of climate-related financial news from the Dow Jones Newswire (2000–2023) it is based on a human-annotated gold standard and a Decompositional Verification Framework (DVF) that decomposes evaluation into completeness, faithfulness, coherence, and relevance, with multi-judge scoring calibrated against human ratings. |